Papers by Joao F. Henriques
HelloFresh: LLM Evalutions on Streams of Real-World Human Editorial Actions across X Community Notes and Wikipedia edits (2024.findings-acl)
Copied to clipboard
Tim Franzmeyer, Aleksandar Shtedritski, Samuel Albanie, Philip Torr, Joao F. Henriques, Jakob Foerster
| Challenge: | a better understanding of LLM capabilities on real world tasks is vital for safe development and deployment. |
| Approach: | They propose a new LLM called HelloFresh that uses real-world data to measure performance . they backtest the model and find it yields a temporally consistent ranking . |
| Outcome: | The proposed benchmarks outperform static evaluation data and test data on Wikipedia pages. |